Standard information retrieval systems often miss deeper meaning when queries are short, vague, or lack specifics. We propose a hybrid framework that pulls structured data from Wikidata and combines dense and sparse retrieval methods to improve both finding documents and generating answers. Our method first identifies key entities in a user query using a lightweight linker, then expands the query with Wikidata properties (instance of, subclass of, part of, different from). A hybrid retriever ColBERTv2 for dense late interaction plus SPLADE for sparse lexical expansion processes the expanded query. The top ten passages then feed a FLAN-T5 generator. A DeBERTa based filter prevents semantic drift and entity confusion. Tests on TREC Deep Learning Track and three BEIR subsets (NQ, HotpotQA, FiQA-2018) show that knowledge-guided expansion outperforms five strong baselines: Standard RAG, HyDE, REPLUG, Self-RAG, and SKR. On TREC DL, nDCG@10 rises from 0.689 to 0.812 (+17.9%); average BEIR nDCG@10 increases from 0.562 to 0.673. BLEU scores for generated answers reach 0.612, a 0.147 absolute gain over Standard RAG (0.465). Ablation tests confirm that adding knowledge graph data provides the largest lift (? nDCG@10 = +0.097). These outcomes show that fusing structured world knowledge with hybrid neural retrieval reduces vocabulary mismatch and improves both retrieval accuracy and answer faithfulness, especially for sparse or ambiguous queries.
Introduction
The text proposes WikKG-RAG (Wikidata Knowledge Graph Retrieval-Augmented Generation), a hybrid information-retrieval and question-answering system designed to improve retrieval and generation for short, vague, or ambiguous queries.
1. depend heavily on exact vocabulary matching. They can therefore miss Problem
Traditional retrieval methods such as TF-IDF and BM25 depend heavily on exact vocabulary matching. They can therefore miss semantically related information when queries are short or ambiguous, such as “COVID treatment” or “memory leak Java.”
Automatic query expansion attempts to solve this vocabulary mismatch, but older statistical approaches can introduce irrelevant terms and cause query drift. Dense retrieval methods such as DPR and ANCE improve semantic matching, but they can still perform poorly when the original query contains too little information.
Although RAG helps ground generated answers in retrieved evidence, poor retrieval can propagate errors into the generation stage.
2. Proposed solution: WikKG-RAG
The proposed system addresses this problem by using Wikidata as a source of structured world knowledge before retrieval.
Instead of simply adding statistically related words, the system selectively adds knowledge-graph entities and relationships that are relevant to the query.
3. Knowledge-guided query expansion
The system extracts important nouns, verbs, and proper nouns from the query and searches Wikidata using four relationships:
P31: instance of
P279: subclass of
P527: part of
P1889: different from
Candidate entities are ranked using three signals:
PMI: measures co-occurrence between the query term and entity.
BGE embedding similarity: measures semantic similarity between the query and entity description.
PageRank: measures the importance of the entity in the knowledge graph.
Only sufficiently relevant entities are added to the query, with a maximum of five expansions per token.
4. Ambiguity handling
The system uses DeBERTa, fine-tuned on MNLI, to determine whether a proposed expansion fits the context of the original query.
This is intended to reduce problems with polysemy, where a word can have multiple meanings. For example, “jaguar” could refer to an animal or a car. Contextual scoring helps select the appropriate interpretation and discard irrelevant expansions.
5. Hybrid retrieval
The enriched query is processed by two complementary retrieval approaches:
ColBERTv2: uses token-level late interaction to identify fine-grained semantic relevance.
The system retrieves the top 10 passages for generation.
This combination aims to obtain the benefits of both semantic retrieval and lexical matching.
6. Answer generation
The retrieved passages and enriched query are passed to FLAN-T5-XXL (11B parameters).
The generator uses the retrieved evidence rather than relying entirely on information stored in its parameters. Generation uses:
Beam size: 5
Maximum output: 256 tokens
This design aims to improve factual grounding and reduce hallucination.
7. Implementation
The experiments were conducted using:
Ubuntu 24.04
Python 3.11
PyTorch 2.3
Two NVIDIA A6000 GPUs
128 GB RAM
Wikidata SPARQL endpoint
Qdrant for entity embeddings
BGE-large for semantic similarity
ColBERTv2 and SPLADE++ for retrieval
FLAN-T5-XXL with 4-bit quantization
The retrieval models were evaluated over the BEIR corpus containing approximately 8.6 million passages.
8. Datasets
The proposed system was evaluated on four benchmarks:
Natural Questions (NQ) – short, real-world search queries that are often ambiguous.
HotpotQA – requires multi-hop reasoning across multiple documents.
FiQA-2018 – financial questions involving domain-specific terminology and entity understanding.
TREC Deep Learning 2023 – passage-ranking benchmark with graded human relevance judgments.
9. Evaluation
The primary reported retrieval metric is nDCG@10, which measures ranking quality while considering both relevance and the position of retrieved documents.
10. Main contribution
The central contribution is the integration of structured knowledge-graph expansion with hybrid dense-sparse retrieval and RAG.
Unlike traditional query expansion, which primarily relies on statistical co-occurrence, WikKG-RAG introduces explicit relationships from Wikidata and uses semantic and contextual filtering to control query drift.
Conclusion
This work shows that injecting structured knowledge measurably improves retrieval?augmented generation. The proposed WikKG?RAG pipeline selectively expands queries with Wikidata entities (instance of, subclass of, part of, different from), filters them with DeBERTa, processes them through a hybrid ColBERTv2/SPLADE retriever, and generates answers with FLAN?T5?XXL. Extensive testing on TREC DL and BEIR confirms effectiveness: nDCG@10 improved 17.9–19.8% over strong baselines, BLEU rose by 0.147 points, and ablations confirm knowledge integration as the primary factor.
References
[1] S. Robertson & H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends in Information Retrieval, 3(4), 2009.
[2] C. Carpineto & G. Romano, “A Survey of Automatic Query Expansion in Information Retrieval,” ACM Computing Surveys, 44(1), 2012.
[3] V. Karpukhin et al., “Dense Passage Retrieval for Open?Domain Question Answering,” EMNLP, 2020.
[4] L. Xiong et al., “Approximate Nearest Neighbor Negative Contrastive Learning for DPR,” ICLR, 2021.
[5] P. Lewis et al., “Retrieval?Augmented Generation for Knowledge?Intensive NLP Tasks,” NeurIPS, 2020.
[6] L. Gao et al., “Precise Zero?Shot Dense Retrieval without Relevance Labels,” arXiv:2212.10496, 2022.
[7] W. Shi et al., “REPLUG: Retrieval?Augmented Black?Box Language Models,” arXiv:2301.12652, 2023.
[8] A. Asai et al., “Self?RAG: Learning to Retrieve, Generate, and Critique through Self?Reflection,” ICLR, 2024.
[9] Y. Wang et al., “Selective Knowledge Retrieval for Augmented Language Models,” ACL, 2024.
[10] D. Vrande?i? & M. Krötzsch, “Wikidata: A Free Collaborative Knowledge Base,” Communications of the ACM, 57(10), 2014.
[11] K. Santhanam et al., “ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction,” NAACL, 2022.
[12] T. Formal et al., “SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval,” SIGIR, 2022.
[13] S. Xiao et al., “C?Pack: Packaged Resources to Advance General Chinese Embedding,” arXiv:2309.07597, 2023.
[14] H. W. Chung et al., “Scaling Instruction?Finetuned Language Models,” arXiv:2210.11416, 2022.
[15] P. He et al., “DeBERTa: Decoding?Enhanced BERT with Disentangled Attention,” ICLR, 2021.
[16] T. Wolf et al., “Transformers: State?of?the?Art Natural Language Processing,” EMNLP?Demo, 2020.
[17] N. Thakur et al., “BEIR: A Heterogeneous Benchmark for Zero?shot Evaluation of Information Retrieval Models,” NeurIPS Datasets, 2021.